# Build plan: the thin golden path Oct 1, 2026 This is the first build: every step of the golden path running end to end on the synthetic study package, thin. It follows `CLAUDE.md` ("build this path thin first, then deepen each step") and changes no PRD scope. Proposals P-2, P-3, P-4 and P-6 from `docs/ceo-review.md` are built in so the owner can judge them on working code. Each can be removed. ## What runs ``` python -m golden_path --seed 7 --out runs/demo ``` ``` HOSPITAL ZONE (process 1, own SQLite file) PLATFORM ZONE (process 2, own SQLite file) ───────────────────────────────────────── ────────────────────────────────────────── synthetic package: hospital/*.csv synthetic package: documents/*.md │ │ intake.deidentify ── de-id log documents.intake ── file types, checklist │ │ store.ClinicalStore (+ consent scope) documents.extract ── protocol fields with │ │ source passages estimator.estimate ── parameters + bootstrap │ estimator.replay ──── cut by date, predict last third ├── study_record.retrospective │ ├── feedback.digest ── themes, citations estimator.export ── gate: cell size ≥ 10, │ │ no patient fields, approver │ ▼ │ exports/exp-XXXX.json ═══ the one crossing ═══▶ profile.store.import_export │ label rule, prediction intervals, │ global defaults (sourced) clinician review (confirm, correct, elicit) │ reports.synopsis ── brief → design │ simulator.engine ×2 ── global vs local simulator.forecast ── enrolment by site type │ reports.design_report ── evidence table, │ sign, export web.static_site ── six pages ``` The runner `golden_path.py` plays the people in both zones. It passes only the export file from one side to the other. ## Decisions taken for this build | Decision | Choice now | Why | Reversal cost | | --- | --- | --- | --- | | Python | 3.12 in a `uv` virtual environment | Already installed; current | None | | Database | SQLite per zone, behind a small store class | PostgreSQL is the target in `CLAUDE.md` but not installed here; the store classes keep the swap local | One module per zone | | Two zones | Two packages, two database files, one export file. A test fails if either package imports the other. | P-4: two servers wait for real data | None | | Web | Static HTML pages built from the design tokens | Front-end framework not chosen | Pages are rendered from plain data; they port to any framework | | Language model | Gateway built and enforced; no hosted provider wired yet | The thin path needs no model; rule-based extraction is the baseline the model must beat | Add a provider class | | Word and PDF | HTML export only | Keeps dependencies small | Add `python-docx` and a PDF renderer | | Minimum cell size | 10 | The PRD does not set one; 10 is common in health statistics | One constant | | Code version | `0.1.0+src.` | No git repository yet | Switch to the commit hash | ## Statistical methods | Quantity | Method | Independent check | | --- | --- | --- | | Dropout hazard by site type | Exponential: events ÷ person-months; patient bootstrap | Hand computation on 5 patients; calibration across seeds | | Dropout curve | Kaplan–Meier at months 1–12 | Hand computation on a small sample | | Visit attendance, screen failure | Proportions; patient bootstrap | Hand computation | | Screening and enrolment rate | Count ÷ site-months active; parametric Poisson bootstrap | Hand computation | | Deviation rate | Count ÷ patient-months | Hand computation | | Outcome SD and observed effect | Sample SD of change; difference in means | numpy directly | | Prediction interval for a new study | Within-study bootstrap SE combined with an elicited between-study spread on the log or logit scale; labelled as such | Formula test | | Power | Simulated t-test from completers per arm | Analytic noncentral-t power within Monte Carlo error | | Enrolment forecast | Poisson processes per site after activation | Back-test in the hospital zone on the ingested study | | Replay | Fit on records dated before the cut; predict enrolments, dropouts and attended visits after it | A test changes every post-cut record and shows the fit does not move | ## Tests (mapped to features) | Test | Feature | Proves | | --- | --- | --- | | `test_boundary` | F9.1 | Neither zone imports the other; the platform never opens the hospital folder | | `test_deidentify` | F9.1 | No name, phone, national ID or address survives; free-text phone numbers are caught | | `test_consent` | F9.2 | The estimator refuses a dataset without a consent scope | | `test_estimators` | F1.5 | Each estimator matches a hand computation | | `test_export_gate` | F1.5, F9.1 | Small cells suppressed; patient fields, missing approver and tampering rejected | | `test_profile_store` | F8.2, F1.6 | `learned` needs 30 patients or 3 sites; reviews need a name; corrections keep the old value; every row validates against the schema | | `test_gateway` | F9.1 | Patient-level input refused and logged | | `test_extract` | F1.2 | Protocol fields correct in English and Vietnamese, each with its passage | | `test_digest` | F1.4 | Theme agreement with the generator's truth; every comment cites its source | | `test_simulator` | F3.6 | Same seed, profile and code version give identical output; power matches the analytic value | | `test_dropout_curve` | F3.4 | Re-running the study's own design reproduces its observed dropout curve | | `test_replay` | F3.8 | No record after the cut reaches the fit | | `test_report` | F2.8, F9.4 | Export blocked without a signer; every evidence row has label, interval and provenance; templates hold no raw numbers | | `test_calibration` | P-3 | Stated 90% intervals hold the truth at close to their stated rate across seeds | | `test_golden_path` | all | The whole path runs and every artefact exists | ## Design plan The six screens in `design/screens/` are the reference. This build renders them as static pages with the same tokens, so the data contract is settled before a framework is chosen. **Components every page uses** | Component | Rule it enforces | | --- | --- | | `LabelMark` | Filled circle `learned`, half circle `elicited`, ring `assumed`, always with the word | | `Quantity` | A value never renders without its mark and interval; the interval type is shown on hover and in the evidence table | | `SampleBanner` | Every page of a synthetic run carries "Sample values. Synthetic study, not real patients." | | `SourceLink` | Every field, theme and parameter links to its passage, export or run | **Changes against the drawn screens** - Report screen: add an interval column and a "spread from" column to the evidence table (P-6). - Profile screen: show the within-study interval and the prediction interval for a new study side by side, so the audience sees how much of the width is elicited. - Simulate screen: show which inputs fell back to a global default, with the default's citation (P-2). **Not yet drawn, needed for the demo:** the new brief (step 7), replay with calibration (step 11), limits (step 12), low-confidence field review, export approval and the audit log. This build renders replay, limits, export approval and the audit log as plain pages. **Language.** Parameter names and report headings carry English and Vietnamese. The report is rendered in both. Full interface translation follows the framework choice. ## After this build 1. Real documents: run backlog spike 0.4 against surrogate Vietnamese protocols; add the model extractor behind the gateway and measure it against the rule baseline. 2. Choose the front-end framework and port the pages. 3. PostgreSQL in each zone. 4. Word and PDF export. 5. Hand-labelled gold sets (backlog Epic 8). ## Results of the first build (1 Oct 2026) The golden path runs end to end in about two seconds and reproduces exactly: a second run with the same seed gives the same export id, the same run ids and the same signed report hash. 71 tests pass. **What the demo now shows, on synthetic data (seed 7)** | Scenario | Enrolled | Simulated power | | --- | --- | --- | | Design A, sized on global defaults, run under global defaults | 330 | 88.2% | | The same design A, run under the local profile | 330 | 72.9% | | Design B, sized on the local profile, run under the local profile | 536 | 87.4% | Design B's simulated power sits below the 90% it was sized for, because the simulation averages over the uncertainty in every input. The report says so. Type I error under the null was 5.15% on 10,000 runs. **Calibration (P-3), 150 synthetic studies** | Interval | Stated | Held the truth | Checks | | --- | --- | --- | --- | | Within-study parameter intervals | 90% | 90.3% | 2,400 | | Replay (dropouts, interim visits) | 80% | 85.8% | 300 | | Enrolment back-test | 80% | 75.3% (95% interval 68% to 82%) | 150 | Replay intervals are a little wide, which errs on the safe side. Counts landing exactly on an endpoint get half credit (mid-p). **Changes the calibration forced** - Parameter intervals moved from percentile bootstrap to a normal interval on the log or logit scale with the bootstrap standard error. Percentile intervals held the truth 88.9% of the time; the new ones hold it 90.3%. - The enrolment back-test cut was defined using the final enrolment date, so where the cut fell depended on the outcome. It now falls on the day half the target was enrolled, a date known at the time. - Replay visit counts were drawn as Poisson, which overstated their spread. Each participant's dropout time is now simulated, and both replay quantities come from the same times. **A leakage bug the replay test caught.** A view dated before the cut showed a visit as missed when it was in fact attended a few days after the cut. A missed visit now counts only once its window has closed on the view date, and an attended visit counts only if it happened by then. **Not done in this build** - Real documents: no scanned-document reading, and no model extraction. The rule-based extractor and theme lexicon are the baselines. Their accuracy here is measured on synthetic text written beside them, so it checks the code, not the method. - PostgreSQL, the job queue, the web framework and Word or PDF export (HTML and Markdown only). - Hand-labelled gold sets (BACKLOG Epic 8). - The clinician review is scripted in `golden_path.py`. There is no review screen with inputs yet. - Numbers on the system pages (seeds, trial counts, calibration rates) are configuration or audit facts and carry no source label. Every number in the report and the profile carries one.